Local-LLM proving ground

Proving local agents before they act.

GAUNTLET Bench runs local models through real workloads — agentic, live, verified — and publishes the evidence.

All scores measured on Apple M5 Max, 128 GB unified memory, LM Studio

Current podium

#1 Generalist — 99 / 100 G Agentic — 100 / 100 A Understanding — 90.1 / 100 U Needle — 89.3 / 100 N Thinking — 96.5 / 100 T Live — 75.7 / 100 L Engineering — 98.5 / 100 E Throughput — 8.1 / 100 T

Qwen3.8-27B Q6_K GGUF

27B dense · Q6_K GGUF

8/8 13.2 tok/s
#2 Generalist — 95.5 / 100 G Agentic — 98 / 100 A Understanding — 79.7 / 100 U Needle — 77.8 / 100 N Thinking — 88 / 100 T Live — 55.5 / 100 L Engineering — 91 / 100 E Throughput — 40.6 / 100 T

Gemma-4 26B-A4B 8bit MLX

26B total / 4B active per token · 8bit MLX

8/8 66.2 tok/s
#3 Generalist — 91 / 100 G Agentic — 100 / 100 A Understanding — 93.2 / 100 U Needle — 99 / 100 N Thinking — 97.2 / 100 T Live — 68.3 / 100 L Engineering — 91.5 / 100 E Throughput — 8.2 / 100 T

Qwen3.6-27B Dense 8bit MLX

27B dense · 8bit MLX

8/8 13.4 tok/s

Leaderboard — 33 models · click a column to sort

# Model Params Quant / Format G A U N T L E T Progress tok/s
1 Qwen3.8-27B Q6_K GGUF 27B dense Q6_K GGUF 9910090.189.396.575.798.58.1 8/8 13.2
2 Gemma-4 26B-A4B 8bit MLX 26B total / 4B active per token 8bit MLX 95.59879.777.88855.59140.6 8/8 66.2
3 Qwen3.6-27B Dense 8bit MLX 27B dense 8bit MLX 9110093.29997.268.391.58.2 8/8 13.4
4 Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX MTPLX 8bit MLX 27B 8bit MLX 95.597.58891.545.51009.4 7/8 15.3
5 Mistral Small 3.2 24B 8bit MLX 24B 8bit MLX 7992.578.780100707.9 7/8 12.9
6 Qwen3.6-35B-A3B 4bit MLX 35B total / 3B active per token 4bit MLX 9387.591.297.193.854.7 6/8 89.4
7 Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX Q8_0 GGUF 27B Q8_0 GGUF 8999.590.597.7896.7 6/8 11
8 Qwen3.6-35B-A3B Unsloth Dynamic UD-Q8_K_XL MLX (Brooooooklyn) 35B total / 3B active per token UD-Q8_K_XL mixed-precision MLX 93.59591.210041.9 5/8 68.4
9 Qwen3.5-122B-A10B 4bit MLX 122B total / 10B active per token 4bit MLX 92.59590.210024.1 5/8 39.4
10 Qwen3-Coder-Next 80B 4bit MLX 80B 4bit MLX 899510071.541.2 5/8 67.3
11 Laguna S 2.1 (Q4_K_M GGUF) 118B total / 8B active (MoE, 10-of-256 experts + 1 shared) Q4_K_M GGUF 869610083.532.6 5/8 53.3
12 Bonsai 27B Ternary 2bit MLX 27B 2bit (ternary, PrismML extreme compression) MLX 82.597.593.19925.2 5/8 41.1
13 Nemotron 3 Nano 30B-A3B 4bit MLX 30B total / ~3B active per token 4bit MLX 79.59272.896.753.1 5/8 86.7
14 Qwen3.8-27B Q4_K_M GGUF 27B dense Q4_K_M GGUF 94.599.595.713.6 4/8 22.2
15 Qwen3.8-27B Q8_0 GGUF 27B dense Q8_0 GGUF 94.599.510014.1 4/8 22.9
16 Gemma-4 E4B 4B 8bit MLX 4B 8bit MLX 82.59310046.5 4/8 76
17 Devstral Small 2 24B 6bit MLX 24B 6bit MLX 78.594.510012.4 4/8 20.3
18 Muse-Glimmer 30B (GGUF, kquant 17GB) 30B (unconfirmed -- filename-derived, see note) kquant (unconfirmed exact scheme -- filename says "kquant", not a standard llama.cpp quant label) GGUF 757710011.1 4/8 18.2
19 GPT-OSS 120B 120B mxfp4 MLX 65.810028.1 3/8 45.9
20 Qwen3-VL 30B-A3B Instruct 4bit MLX 30B total / 3B active per token 4bit MLX 81.710030.2 3/8 49.2
21 Qwen3.6-35B-A3B MLX 8bit Uniform (lmstudio-community) 35B total / 3B active per token 8bit uniform MLX 85.5n/d53.6 2/8 87.5
22 Gemma-4 E4B 7.5B Q4_K_M GGUF 7.5B Q4_K_M GGUF 78n/d54.7 2/8 89.4
23 Devstral Small 2 24B 4bit MLX 24B 4bit MLX 74.5n/d21.2 2/8 34.7
24 Hermes-4 70B 4bit MLX 70B 4bit MLX 73n/d6.4 2/8 10.4
25 Qwen3-4B-2507 (non-thinking) 4bit MLX 4B 4bit MLX 68.5n/d84.3 2/8 137.6
26 GPT-OSS 120B Fable-5 Distilled 120B mxfp4 MLX 62.7n/d21.3 2/8 34.8
27 GPT-OSS 20B (OpenAI mxfp4) 20B mxfp4 MLX 62.3n/d53.3 2/8 87.1
28 Gemma 3 4B QAT 4bit MLX 4B QAT 4bit MLX 61.5n/d100 2/8 163.3
29 DeepSeek-R1-Distill-Qwen-32B 8bit MLX 32B 8bit MLX 60.5n/d8 2/8 13.1
30 DeepSeek-R1-Distill-Llama 70B 8bit MLX 70B 8bit MLX 58n/d3.5 2/8 5.7
31 GPT-OSS 20B MLX 20B MLX 52.3n/d20.7 2/8 33.8
32 GPT-OSS Safeguard 20B MLX 20B MLX 50.9n/d35.5 2/8 58
33 Phi-3.5-mini-instruct 4bit MLX 3.8B 4bit MLX 31.5n/d67.5 2/8 110.2

Axes are 0–100. = suite not yet run for that model. n/d = Thinking axis needs ≥15 observed tests to score (see methodology). Default order: GAUNTLET progress, then Generalist score.